Papers by Mohammed Safi Ur Rahman Khan
MILU: A Multi-task Indic Language Understanding Benchmark (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing benchmarks focus on English, leaving substantial gaps in assessing LLM capabilities in low-resource and linguistically diverse languages. |
| Approach: | They propose a multi-task indic language understanding benchmark to assess LLMs in low-resource languages. |
| Outcome: | The new benchmark spans 8 domains and 41 subjects across 11 Indic languages, reflecting general and culturally specific knowledge. |
Data and Model Centric Approaches for Expansion of Large Language Models to New languages (2025.emnlp-tutorials)
Copied to clipboard
| Challenge: | Existing LLMs mainly support English alongside a handful of high resource languages . this leaves a major gap for most low-resource languages despite increasing pace of research . |
| Approach: | This tutorial examines approaches to expand the language coverage of LLMs . they look at tokenizer training, pre-training, instruction tuning, alignment, evaluation, etc. |
| Outcome: | This tutorial examines approaches to expand the language coverage of LLMs . it provides an efficient and viable path to bring LLM technologies to low-resource languages . |
FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes (2025.acl-long)
Copied to clipboard
Janki Atul Nawale, Mohammed Safi Ur Rahman Khan, Janani D, Mansi Gupta, Danish Pruthi, Mitesh M Khapra
| Challenge: | Existing studies on fairness of LLMs are largely Western-focused, making them inadequate for culturally diverse countries such as India. |
| Approach: | They propose a benchmark to evaluate fairness of LLMs across 85 identity groups . they consult domain experts to curate over 1,800 socio-cultural topics . |
| Outcome: | The benchmark evaluates LLMs across 85 identities across 85 castes, religions, regions, and tribes. |
Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian Languages (2025.acl-long)
Copied to clipboard
Ashwin Sankar, Sparsh Jain, Nikhil Narasimhan, Devilal Choudhary, Dhairya Suman, Mohammed Safi Ur Rahman Khan, Anoop Kunchukuttan, Mitesh M Khapra, Raj Dabre
| Challenge: | Existing datasets that cover only a fraction of Indian languages lack the breadth needed to generalize beyond curated benchmarks. |
| Approach: | They propose to build the largest speech translation dataset for Indian languages . they use a three-step methodology to gather data and train a model that performs better . |
| Outcome: | The proposed model improves on existing models and is open-source with permissive licenses. |
Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs (2025.acl-long)
Copied to clipboard
Sumanth Doddapaneni, Mohammed Safi Ur Rahman Khan, Dilip Venkatesh, Raj Dabre, Anoop Kunchukuttan, Mitesh M Khapra
| Challenge: | Evaluating machine-generated text remains a challenge in NLP for non-English languages . current evaluation frameworks focus on English, revealing a gap in multilingual evaluations . |
| Approach: | They propose a cross-lingual auto evaluation framework that includes evaluator LLMs and a test set specifically designed for multilingual evaluation. |
| Outcome: | The proposed model aligns more closely with human judgments than proprietary models on non-English language evaluations. |
Can Vision-Language Models Evaluate Handwritten Math? (2025.acl-long)
Copied to clipboard
| Challenge: | Recent advances in Vision-Language Models (VLMs) have significantly enhanced the ability to interpret both textual and visual data. |
| Approach: | They propose a benchmark to assess VLMs’ ability to detect, localize and correct errors in handwritten mathematical content. |
| Outcome: | The proposed benchmark covers over 2,200 handwritten math solutions from 609 manually curated problems from grades 7-12 with intentionally introduced perturbations. |